Papers with Elo ratings

3 papers
Style Over Substance: Evaluation Biases for Large Language Models (2025.coling-main)

Copied to clipboard

Challenge: Ranking the relative performance of large language models based on Elo ratings is gaining popularity . however, the extent to which humans and LLMs are capable evaluators remains uncertain .
Approach: They propose to evaluate machine-generated text across multiple dimensions using the Elo rating system . they propose to use crowd-sourced and expert annotators to rank models based on Elo ratings .
Outcome: The proposed method improves the quality of LLM-based evaluations, but there is no improvement in crowd-sourced evaluations.
War of Thoughts: Competition Stimulates Stronger Reasoning in Large Language Models (2025.findings-acl)

Copied to clipboard

Challenge: Recent advances in Large Language Models (LLMs) have reshaped the landscape of reasoning tasks.
Approach: They propose a method that enhances LLM reasoning without finetuning by using test-time scaling.
Outcome: The proposed method outperforms baseline models in both budget and model size.
PsychePass: Calibrating LLM Therapeutic Competence via Trajectory-Anchored Tournaments (2026.findings-acl)

Copied to clipboard

Challenge: evaluating therapeutic competence of large language models remains challenging due to unstructured and longitudinal nature of counseling.
Approach: They propose a framework that calibrates the therapeutic competence of LLMs via trajectory-anchored tournaments.
Outcome: The proposed framework calibrates the therapeutic competence of LLMs via trajectory-anchored tournaments.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations